Goto

Collaborating Authors

 person view




Grounded Gesture Generation: Language, Motion, and Space

arXiv.org Artificial Intelligence

Human motion generation has advanced rapidly in recent years, yet the critical problem of creating spatially grounded, context-aware gestures has been largely overlooked. Existing models typically specialize either in descriptive motion generation, such as locomotion and object interaction, or in isolated co-speech gesture synthesis aligned with utterance semantics. However, both lines of work often treat motion and environmental grounding separately, limiting advances toward embodied, communicative agents. T o address this gap, our work introduces a multi-modal dataset and framework for grounded gesture generation, combining two key resources: (1) a synthetic dataset of spatially grounded referential gestures, and (2) MM-Conv, a VR-based dataset capturing two-party dialogues. T ogether, they provide over 7.7 hours of synchronized motion, speech, and 3D scene information, standardized in the HumanML3D format. Our framework further connects to a physics-based simulator, enabling synthetic data generation and situated evaluation. By bridging gesture modeling and spatial grounding, our contribution establishes a foundation for advancing research in situated gesture generation and grounded multimodal interaction.


Pipeline for 3D reconstruction of the human body from AR/VR headset mounted egocentric cameras

arXiv.org Artificial Intelligence

In this paper, we propose a novel pipeline for the 3D reconstruction of the full body from egocentric viewpoints. 3-D reconstruction of the human body from egocentric viewpoints is a challenging task as the view is skewed and the body parts farther from the cameras are occluded. One such example is the view from cameras installed below VR headsets. To achieve this task, we first make use of conditional GANs to translate the egocentric views to full body third-person views. This increases the comprehensibility of the image and caters to occlusions. The generated third-person view is further sent through the 3D reconstruction module that generates a 3D mesh of the body. We also train a network that can take the third person full-body view of the subject and generate the texture maps for applying on the mesh. The generated mesh has fairly realistic body proportions and is fully rigged allowing for further applications such as real-time animation and pose transfer in games. This approach can be key to a new domain of mobile human telepresence.


Shaping Belief States with Generative Environment Models for RL

arXiv.org Artificial Intelligence

When agents interact with a complex environment, they must form and maintain beliefs about the relevant aspects of that environment. We propose a way to efficiently train expressive generative models in complex environments. We show that a predictive algorithm with an expressive generative model can form stable belief-states in visually rich and dynamic 3D environments. More precisely, we show that the learned representation captures the layout of the environment as well as the position and orientation of the agent. Our experiments show that the model substantially improves data-efficiency on a number of reinforcement learning (RL) tasks compared with strong model-free baseline agents. We find that predicting multiple steps into the future (overshooting), in combination with an expressive generative model, is critical for stable representations to emerge. In practice, using expressive generative models in RL is computationally expensive and we propose a scheme to reduce this computational burden, allowing us to build agents that are competitive with model-free baselines.


Video of Wembley Stadium hosting Drone Racing League in London

Daily Mail - Science & tech

The Drone Racing League flew in to the capital this week as machines soared around Wembley Stadium at speeds of 75mph (120km/h). Drones buzzed around the iconic venue and further proved why the sport of drone racing is gaining popularity. The racing was live streamed to spectators for the first time over EE's 4G network at the stadium, with 4G cameras attached to the drones giving people a drone's-eye-view. The event was attended by 16-year-old Luke Bannister, from Somerset, who recently won 174,000 ( 250,000) in the Drone Grand Prix in Dubai. First Person View (FPV) drone racing involves live video being streamed to the pilot's headset to enable split-second manoeuvres. This perspective is usually only available to the team controlling the drone, however for the first time spectators in the stadium and online were also able to'ride' around the stadium.